Back

Computational and Structural Biotechnology Journal

American Association for the Advancement of Science (AAAS)

Preprints posted in the last 7 days, ranked by how well they match Computational and Structural Biotechnology Journal's content profile, based on 242 papers previously published here. The average preprint has a 0.23% match score for this journal, so anything above that is already an above-average fit.

1
Geometric characterization of the HSV - 1 glycoprotein B - amyloid β interaction in Alzheimer's disease using Forman-Ricci curvature

Bou Dagher, L.; Han, Z.; Zhou, S.; Fülöp, T.; Desroches, M.; Rodrigues, S.

2026-08-29 bioinformatics 10.64898/2026.08.26.747308 medRxiv
Top 0.1%
12.2%
Show abstract

Alzheimer's disease is characterized by the accumulation and aggregation of amyloid-{beta}(A{beta}), but the molecular mechanisms linking environmental and infectious factors to A$\beta$ conformational changes remain incompletely understood. Herpes simplex virus type 1 (HSV-1) has been proposed as a potential contributor to AD pathology, and interactions between the viral glycoprotein B (gB) and A$\beta$ may influence the conformational behaviour of the peptide. Molecular dynamics (MD) simulations provide atomic-scale information on such interactions, but conventional structural descriptors may not fully capture changes in the organization of residue interaction networks. Here, we introduce a graph-geometric framework based on Forman-Ricci curvature to characterize the evolution of residue interaction networks during MD simulations. Each simulation frame is represented as a residue interaction graph based on C--C contacts, and residue-wise curvature profiles are analysed across time. We apply the framework to A{beta}1-42 in isolation and in complex with HSV-1 gB. Conventional MD analyses indicate stable association of the simulated complex, favourable interaction energetics, and conformational changes in A{beta}, including a transition from -helical structure toward {beta}-turn-rich conformations over the simulated timescale. Forman-Ricci curvature reveals pronounced and spatially localized remodelling of the A{beta} residue interaction network in the complex, with the strongest changes concentrated in the C-terminal region. These regions also exhibit reduced temporal curvature fluctuations and progressively distinct geometric behaviour throughout the simulation. Hierarchical clustering further identifies cooperative groups of residues with coordinated curvature dynamics, including a prominent C-terminal domain. Together, these results demonstrate that Forman-Ricci curvature provides a complementary description of biomolecular dynamics by capturing changes in the geometric organization of residue interaction networks that are not directly represented by conventional structural descriptors. The framework provides a general computational approach for studying network-level structural remodelling in protein molecular dynamics and offers a quantitative perspective on the conformational consequences of HSV-1 gB--A{beta} association.

2
MOSurvivor-Guided Joint CpG Selection and XGBoost Hyperparameter Optimization for Compact Epigenetic Age Prediction

Yelgi, A.; Tavangari, S.; Shakarami, Z.; Janfaza, S.

2026-08-29 genomics 10.64898/2026.08.26.747213 medRxiv
Top 0.1%
9.0%
Show abstract

Accurate epigenetic age prediction from DNA methylation profiles is intrinsically high-dimensional, creating a need for parsimonious models that preserve predictive performance while reducing the number of assayed cytosine-phosphate-guanine (CpG) loci. This study introduces MOSurvivor, a population-based multi-objective search framework that jointly optimizes a weight-threshold CpG selector and eight XGBoost hyperparameters. Experiments used the GSE40279 whole-blood cohort (656 individuals profiled on the Illumina HumanMethylation450 platform). After retaining 1,000 age-correlated CpGs, five strategies were evaluated on the same 30 seeded 80:20 train/test splits: fixed-parameter XGBoost using all 1,000 CpGs, random search, a genetic algorithm, particle swarm optimization, and MOSurvivor. Internal fitness was estimated using three-fold cross-validation on each training set. Across the 30 held-out test sets, MOSurvivor achieved a mean absolute error (MAE) of 4.149 {+/-} 0.300 years, root mean squared error of 5.545 {+/-} 0.392 years, and R2 of 0.855{+/-} 0.027 while retaining 211.6 {+/-} 54.8 CpGs. Relative to full-feature XGBoost (MAE 4.095 {+/-} 0.285 years), MOSurvivor reduced the feature set by 78.8% at an MAE increase of only 0.054 years (1.3%). Paired Wilcoxon tests found no significant accuracy difference between MOSurvivor and any comparator (all unadjusted p > 0.05; all Holm-adjusted p [≥] 0.476). The most recurrent locus, cg16867657, appeared in 29 runs, whereas mean pairwise Jaccard similarity was 0.124, indicating a small stable core embedded in multiple near-equivalent feature subsets. MOSurvivor thus offers a competitive accuracy-parsimony trade-off rather than superior absolute accuracy. External validation and leakage-free nested feature preselection remain necessary before biological or clinical translation. Keywords: epigenetic clock, DNA methylation, feature selection, multi-objective optimization, XGBoost, metaheuristics, biological aging.

3
Making Accelerating Medicines Partnership Data Findable and Interoperable through a Common Data Model: Extending OMOP for Multi-Source Multimodal Data

Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.

2026-09-02 genetic and genomic medicine 10.64898/2026.08.31.26361831 medRxiv
Top 0.3%
6.7%
Show abstract

SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.

4
A Simple Method to Distinguish Active and Inactive Aptamers by Analyzing the Ruggedness of the Aptamer Free Energy Landscape

Subramanian, G.; Thiel, W.; Singh, R.

2026-08-29 bioinformatics 10.64898/2026.08.26.747184 medRxiv
Top 0.3%
6.7%
Show abstract

Aptamers are structured nucleic acid ligands capable of high affinity, high specificity molecular recognition generated using variations of the SELEX (Systematic Evolution of Ligands by Exponential Enrichment) process. However, SELEX often produces sequences that enrich yet may lack binding efficacy. We propose a measure called the Ruggedness Composite Index (RCI) along with a method for computing it, that can be used to distinguish binding-competent ('active') aptamers from weak or non-binding ('inactive') aptamers. Given a set of aptamers, RCI incorporates information on their fragmentation (landscape partitioning), basin entropy (metastable state distribution), cumulative density irregularity (non-uniform occupancy), and structural energy correlation length (structure-energy coupling scale). We test whether secondary-structure folding energy landscape topology distinguishes active from inactive aptamers using a multiscale level set framework across six datasets. Active aptamers show lower RCI values and occupy smoother, funnel-like conformational spaces, while inactive aptamers show higher RCI values, reflecting fragmented, high-entropy landscapes. By contrast, classical thermodynamic features, such as minimum free energy, show limited discrimination between active and inactive aptamers. In all datasets, sequences that exhibit enrichment which is not monotonic but lack specificity exhibit elevated ruggedness, indicating landscape topology can predict non-specific enrichment. These results indicate that folding landscape organization can be used as a predictor of aptamer activity and establish RCI as a simple, mechanistically interpretable measure for improving candidate prioritization, especially in therapeutic aptamer discovery.

5
Prot2Surf: fast analysis of protein - surface binding modes

Muniz-Chicharro, A.; Tanriver, G.; Gora, A.

2026-08-29 bioinformatics 10.64898/2026.08.26.747352 medRxiv
Top 0.8%
4.5%
Show abstract

Summary: Prot2Surf is a software tool designed for the characterization and prediction of protein association to surfaces. In this application note, Prot2Surf was tested using catalytic domains of the lytic polysaccharide monooxygenases (LPMOs), interacting with native surfaces. The results show that the software can efficiently analyze key binding features, including protein-surface distances, distances between catalytically reactive atoms, and the orientation angle between surface chains and the protein. These features are essential for distinguishing productive binding poses in these protein-surface systems and for understanding interaction patterns that provide guidance on protein engineering. Prot2Surf performs these analyses within seconds to a few minutes, providing a fast and accessible framework to post-process and characterize protein-surface encounter complexes. Availability and implementation: Prot2Surf, which is written in Fortran90, is documented and freely available as open source on GitHub: https://github.com/TUNNELING-GROUP/Prot2Surf. In order to run Prot2Surf, users should also install the SDA software package which is freely available at https://www.h-its.org/downloads/sda7/.

6
AURORA: Analysing and understanding responses to oncological regimens with artificial intelligence

Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.

2026-09-02 health informatics 10.64898/2026.08.30.26361778 medRxiv
Top 2%
2.4%
Show abstract

Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.

7
A Scalable Biological Clock for Metabolic Disease Prediction from the Phenome India Cohort

Tiwari, P.; Garg, M.; Pattanayak, S.; Sarkar, I.; Roy, R.; Bhatraju, N.; Verma, A.; K, S. R.; Prakash, S.; Kumar, V. S.; Uddin, M. A.; Rawat, N.; Sahu, A.; Kumar, Y.; Leuva, P. H.; Mridha, A.; Yenamandra, V.; Singh, A. P.; Mishra, A.; Raychaudhuri, S.; Tallapaka, K. B.; Chandak, G. R.; Kulkarni, M. J.; Dharne, M.; Wahengbam, R.; Kalita, J.; Manna, P.; Subudhi, U.; Majumder, S.; Chakraborty, P.; Chaudhary, K.; Sengupta, S.; Phenome India Consortium, ; Sardana, V.; Chatterjee, S.; Ganguly, D.

2026-09-03 endocrinology 10.64898/2026.08.29.26361656 medRxiv
Top 3%
1.8%
Show abstract

Background: India has a rising incidence of chronic non-communicable diseases, making it a major healthcare burden today. Growing evidence suggests that chronic low-grade inflammation links ageing with cardiometabolic disorders, captured by the emerging concept of inflammaging. However, most evidence on biological ageing comes from Western populations, with no similar models developed for the Indian population. Given the country's distinctive genetic makeup, unique exposome, and heterogeneous NCD presentation, Western models may not capture inflammaging and its effects in the Indian population. Methods: We analysed baseline data from 4,240 adults in the Phenome India CSIR Health Cohort Knowledgebase (PI CheCK), a nationwide multi-centre cohort. Participants were stratified into eight cardiometabolic phenotype groups by BMI (Asian cut off), blood pressure and HbA1c status. We trained a Super Learner ensemble to predict chronological age in the lean normotensive-normoglycaemic reference group (n=615) using 44 plasma cytokines, sex, haemoglobin, and bioimpedance-derived visceral fat area, per cent body fat, and total body water. Performance was assessed by repeated five-fold cross-validation and in a held-out healthy test set. Calibrated biological age acceleration was then estimated in the remaining 3,625 participants. Results: Median age was 51.0 years (IQR 41.0 to 62.0) and 49.4% were female. The Super Learner outperformed elastic net and XGBoost comparators. Permutation importance identified visceral fat area, per cent body fat, CTACK, SDF1a, haemoglobin and sex as leading contributors, with body composition measures accounting for the largest share, indicating an immune-metabolic rather than cytokine-only signal. Biological age acceleration was concentrated in overweight/obese phenotypes. Lean phenotypes showed acceleration close to the reference (0.32 0.50 years). Conclusions: Cytokine and body composition measures capture a quantifiable immunometabolic ageing signal in a South Asian cohort, with acceleration driven predominantly by adiposity. External validation and longitudinal follow up are required.

8
External Validation of a Mathematical Model of Brain Health

Sadia, H.; Doyon, N.; Duchesne, S.

2026-09-03 neurology 10.64898/2026.09.01.26361929 medRxiv
Top 3%
1.8%
Show abstract

Background Understanding the mechanisms underlying brain aging and age-related pathological changes is essential for advancing brain health research. Our group previously developed a mechanistic mathematical model of healthy brain, Chamberland et al. (2024) that integrates key biological processes involved in normal aging, from which Alzheimer's disease (AD) related changes may emerge naturally. Objectives To characterize and validate this brain model by evaluating its sensitivity, calibrating its parameters, and assessing generalizability in independent populations. Methods The model represents the evolution of key biological processes associated with brain aging, including amyloid beta (A{beta}), tau pathologies, neuroinflammation, and neuronal death. After identifying the 30 most influential parameters, we calibrated the model using cognitively normal (CN) participants from the AD Neuroimaging Initiative (ADNI) database (n = 211) by minimizing a loss function composed of three outcomes (AB) plaques, tau tangles, and neuronal density). The calibrated model was then applied to the UK Biobank cohort (n = 35,899) of normal controls (aged 44-82 years). The effects of sex and APOE were evaluated using stratified simulations. Results Parameter calibration significantly reduced the prediction errors for A{beta} and tau. Neuronal density predictions showed strong agreement in the UK Biobank cohort. The variance decomposition identified APOE status as a major contributor to variability in A{beta}. Conclusion Our validated brain health model links mechanistic pathways with population data and reproduces neuronal density patterns in an independent cohort. These findings support its use as a framework for studying brain aging and investigating how Alzheimer's disease related pathological changes may emerge with aging.

9
Towards transferable explicit-solvent coarse-grained models for biomolecular condensates

Toplek, F. B.; Borges-Araujo, L.; Lindorff-Larsen, K.; Everaers, R.; Souza, P. C. T.; Morozova, T. I.

2026-08-29 biophysics 10.64898/2026.08.27.747511 medRxiv
Top 4%
1.7%
Show abstract

Biomolecular condensates formed by intrinsically disordered proteins require molecular models that accurately describe proteins in both dilute solution and condensed phases. Explicit-solvent coarse-grained models offer an attractive balance between chemical resolution and computational efficiency. Yet, it remains unclear whether improving dilute-state properties is sufficient to obtain an accurate description of condensates. Here, we address this question by introducing minimal modifications to the Martini 3 force field that combine recent advances in bonded interactions with refined protein-water interactions and strengthened glycine self-interactions, while preserving the underlying chemical transferability of the model. The resulting model substantially improves the description of single-chain conformations across a diverse benchmark of disordered proteins. We then investigate phase separation of the well-characterized low-complexity domain of heterogeneous nuclear ribonucleoprotein A1 and its sequence variants. The model reproduces several key physicochemical properties of biomolecular condensates, including chain expansion in the dense phase, sequence-dependent intermolecular contacts, protein diffusion and its relation to single-chain dimensions, and hydration, while revealing quantitative limitations in condensate density, phase equilibria, and ion partitioning. Our results show that improving dilute-state behaviour translates into a better description of condensed-phase properties, including condensate density, but is not sufficient to quantitatively reproduce the equilibrium between the dilute and dense phases.

10
Machine Learning-Based Prediction of Maternal Morbidity across Heterogeneous Populations in the United States using Sequential Modeling of the All of Us Dataset

Zhuang, H.; Zakama, A.; Heller, K.; Faulkner, S.; Gollub, B.; Young-Lin, N.; Chen, I. Y.; Asiedu, M.

2026-08-31 obstetrics and gynecology 10.64898/2026.08.25.26360552 medRxiv
Top 4%
1.5%
Show abstract

In this work, we demonstrate the unprecedented value of NIH's "All of Us Research Program" (AoURP) dataset in studying maternal morbidity and building predictive machine learning (ML) models across heterogeneous populations in the United States. We developed robust and data-driven preprocessing pipelines to curate a longitudinal, multi-site, multimodal, and demographically diverse pregnancy dataset (20,253 subjects; 27,525 pregnancy episodes) from AoURP data, using electronic health records (EHR) (Conditions, Labs, Measurements) and survey responses (Social Determinant of Health (SDoH)), focusing on 7 crucial maternal health adverse outcomes. After characterizing data quality, missingness, and heterogeneity, we performed statistical correlation analysis to identify risk factors. We subsequently developed XGBoost and sequential LSTM models to predict the adverse outcomes, reaching state-of-the-art performance for multiple outcomes. We conducted model interpretability post-hoc analysis to understand success points and fairness analysis to evaluate implications for socio-economic disparities. Four practicing physicians reviewed the set of statistically significant and ML model identified features to assess their clinical validity and novelty. Most features identified through either statistical correlations or ML feature importance analysis aligned with known clinical risk factors. Several features were identified that the ML models used but that are not currently used in clinical practice and may merit further clinical investigation. Fairness analysis revealed certain associations with SDoH and age highlight areas that warrant continued monitoring. Overall, we demonstrate that meaningful populational level patterns can be extracted, and high-performing machine learning models can be trained on this longitudinal, diverse, multi-site dataset. Important risk features, particularly novel ones identified, if validated, could inform new strategies for maternal care or enable development and validation of outcome-specific, clinically deployable ML models.

11
Software Application Profile: A real-time surveillance system for monitoring heat exposure and its health impacts - presenting the Rio de Janeiro Heat Dashboard

de Araujo Morais, J. H.; Dias Ferreira, C.; Saraceni, V.; Medeiros de Oliveira Cruz, D.; Mateus Oliveira Aguilar, G.; Cruz, O. G.

2026-08-31 epidemiology 10.64898/2026.08.26.26361449 medRxiv
Top 4%
1.2%
Show abstract

Motivation: With the scaling frequency and intensity of extreme heat events across the globe, it is critical for public institutions to develop early detection systems and continuous monitoring of these events and their impacts. In Brazil, Rio de Janeiro was the first city to publish its heat protocol, with the Rio Heat Dashboard as a central component of this system. Implementation: The dashboard was implemented using R/Shiny and integrates climatic and health data from multiple sources. General features: The application comprises real-time heat exposure monitoring and automatic alert level classification, which is monitored daily by multiple municipal actors and supports activation of actions specified in the heat protocol. It also features a health impact module, which lists each heat event and its impact on mortality, and primary care and emergency visits. Availability: The source for full reproducibility is available through https://github.com/joaohmorais/RioHeatDashboard.

12
Adverse drug withdrawal event signals in FAERS and Eudravigilance databases: a stratified disproportionality analysis study

Khan, Z.; McCarthy, C.; Dalton, K.; Jungo, K. T.; Doherty, A. S.; Reeve, E.; Moriarty, F.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.29.26361707 medRxiv
Top 5%
1.1%
Show abstract

Background: Adverse drug withdrawal events (ADWEs) are a key safety concern during deprescribing but remain poorly explored in pharmacovigilance systems. Objectives: To identify and compare ADWE signals across drug classes, different drugs within drug classes, and across patient characteristics, countries, and over time. Methods: A case/non-case disproportionality analysis was conducted in FDA-FAERS and EMA-EudraVigilance pharmacovigilance databases, with stratification by age (adults: 18-64, older adults: [&ge;]65), sex (male/female), reporting time (2004-2023 in 5-year intervals), and country (for EMA data). Disproportionality analysis (quantitative signal detection) was used to detect signals between ADWEs and drugs using the proportional reporting rate (PRR[&ge;]2), reporting odds ratio (ROR>1), and information component (IC>0) with case count [&ge;]5. Results: Overall, 158,501 reports (FDA-FAERS 145,514; EMA-EudraVigilance 12,987) included drug-event pairs related to ADWEs. In FDA-FAERS, clobetasone (IC=5.58; PRR=79.18; ROR=176.90) showed the strongest ADWE signals, followed by hydromorphone (4.85; 29.94; 37.37), hydrocodone, and paroxetine. In EMA-EudraVigilance, ethyl loflazepate (IC=6.01; PRR=119.80; ROR=197.53), clobetasone (5.39; 102.73; 155.10), veralipride, and levomethadone had the strongest signals. Most drugs maintained positive ADWE signals in analysis stratified into adults and older adults. However, among the top 10 drugs (based on highest IC values), buprenorphine/naloxone, desvenlafaxine, and baclofen in FDA-FAERS (ICs 4.95-6.05) showed stronger signals in older adults. A sex-based difference was observed, with paroxetine, venlafaxine, and buprenorphine/naloxone showing a stronger positive signal in females in both databases, whereas several opioids had stronger signals in males versus females across both databases. Conclusion: This study suggests ADWE signals for some medications differ by age and sex, potentially indicating different risks for withdrawal effects.

13
Limits of Trial-Adaptive Neural Language Fusion Across Large Language Models in P300 Brain Computer Interfaces

Gorenshtein, A.; Omar, M.; Jia, E. L.; Adiniaev, Y.; Daniel, O.; Kruskal, J.; Ahmed, M.; Brook, O. R.; Klang, E.; Barash, Y.

2026-09-03 neurology 10.64898/2026.08.30.26361777 medRxiv
Top 5%
1.1%
Show abstract

Objective: Published P300-speller fusion schemes fix prior trust regardless of trial reliability; we tested whether a reliability estimate improves on it. Methods: We reanalyzed 3,373 archived P300-speller selections from 47 people with ALS (BigP3BCI). A fair, matched-search-space comparison, tuning both a fixed weight and an adaptive policy out-of-fold, was evaluated across 22 evaluable language-model priors up to 46.7B parameters. Two representative priors, GPT-2 and a classical 5-gram, additionally received detailed naive and mechanistic analyses. Results: No prior's 95% CI favored adaptive fusion under the fair comparison, despite unexploited oracle headroom at every scale. Under GPT-2, the naive comparison was significantly worse for adaptive fusion; both anchors converged to a degenerate or near-degenerate fair-comparison solution. For the representative anchors, three further controllers failed to convert that headroom into benefit; the fixed-fused posterior's output probability outperformed the best controller for flagging errors (2.8- to 3.8-fold enrichment). Conclusion: A tuned fixed weight is a difficult-to-beat default across the tested scale range; reliability estimation gave no deployable adaptive advantage. Significance: Adaptive weighting should be validated against a fairly tuned baseline across model families and scales; in this dataset, the fused output's confidence identified high-risk selections better than the tested purpose-built ranker.

14
A germline KDM3C polymorphism impairs DNA repair and sensitizes to chemoradiotherapy

Hasan, A.; Demidova, E. V.; Priyadarshini, P.; Czyzewicz, P.; Gathuka, L.; Murayama, T.; Zhou, Y.; Kiss, Z. A.; Shastry, R. K.; Andrake, M.; Hearne, G.; Devarajan, K.; Wu, C.; Shah, A.; Schultz, B. M.; Connolly, D. C.; Rosen, G. L.; Canadas, I.; Liu, J. C.; Burtness, B. A.; Smith, J. J.; Dunbrack, R. L.; Golemis, E. A.; Whetstine, J. R.; Meyer, J. E.; Arora, S.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.26.26360896 medRxiv
Top 5%
1.0%
Show abstract

Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.

15
SALRR: Scalable Analysis of Long-Read RNA-Seq Enables Comprehensive Transcriptome Profiling in Human Brain

Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.

2026-08-29 genomics 10.64898/2026.08.27.747499 medRxiv
Top 6%
1.0%
Show abstract

Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.

16
Hormonal Therapies For Endometriosis: A Systematic Review And Meta-Analysis Of Randomised Head-To-Head Trials

Bandini, V.; Whitaker, L. H.; Vincent, K.; Salmeri, N.; Mawson, R.; Vercellini, P.; Horne, A. W.

2026-08-31 obstetrics and gynecology 10.64898/2026.08.26.26361442 medRxiv
Top 6%
0.9%
Show abstract

Background: Endometriosis is a chronic pain condition in which hormonal therapies form the cornerstone of long-term management. Treatment tolerability is critical for adherence and therapeutic success, but most comparative studies and reviews have focused on their ability to reduce menstrual pain, while their impact on non-menstrual pelvic pain (NMPP), bleeding patterns, adverse events (AEs), treatment discontinuation and quality of life (QoL) remain poorly characterised. This systematic review and meta-analysis evaluate these outcomes across currently available hormonal therapies, providing practical evidence for clinical decision-making. Methods: PubMed/MEDLINE, Scopus, and Embase were searched up to November 2025 for randomised controlled trials comparing at least two active first- or second-line hormonal treatments for endometriosis. Studies without confirmed endometriosis, treatment duration less than three months and comparing therapies to placebo only were excluded. Data were extracted by two reviewers from reports. Pain outcomes were pooled as mean differences (MD, 95% CI), with bleeding patterns, AEs, and discontinuations as proportions. Analyses were performed in R. PROSPERO: CRD420251137785. Findings: Of 1892 records screened, 48 trials (5583 women) met our inclusion criteria. Overall pelvic pain (0-10 scale) was significantly reduced across all treatment categories (p<0.001): combined oral contraceptives (COCs) (MD 3.17), oral and long-acting progestogens (MD 3.83; MD 4.29), and GnRH-analogues (MD 3.81). Sensitivity analyses restricted to studies reporting NMPP yielded comparable results. GnRH-agonists showed the most favourable bleeding profile, followed by continuous COCs. However, all regimens reported class-specific AEs, including mood changes, nausea, headache, weight gain, and decreased libido (pooled proportions >10%). Overall discontinuation due to AEs was 7.7%, and vaginal bleeding was the leading cause. Heterogeneity across meta-analyses was high. Risk of bias (RoB2) was moderate to high. Interpretation: Given similar reductions in overall pelvic pain across hormonal therapies, treatment decisions should prioritise differences in bleeding profiles, therapy-specific AEs, and QoL. Funding: None.

17
nf_xpatial: A Reproducible Framework for Standardized Preprocessing and Clustering of Xenium Data

Potter, L. A.; Trull, A.; Kumar, N.; Drake, O. R.; Nogueira, M.; Peters, J.; Heinsbroek, J. A.; Day, J. J.; Worthey, E. A.; Ianov, L.

2026-08-29 bioinformatics 10.64898/2026.08.25.747147 medRxiv
Top 8%
0.6%
Show abstract

Recent advances in spatial transcriptomics have enabled the profiling of increasingly larger numbers of genes while retaining single-cell and subcellular resolution in situ. However, standardized bioinformatics workflows for analyzing these datasets have lagged behind, with existing pipelines focusing primarily on image processing and cell segmentation. To address this gap, we present nf_xpatial, a best-practices Nextflow pipeline for the downstream analysis of 10x Genomics Xenium data. The pipeline performs quality control, filtering, log and cell area normalization, multi-sample integration, and both expression-driven and spatially informed clustering across systematic parameter sweeps, allowing users to evaluate and compare clustering resolutions and spatial modeling parameters within a single reproducible run. Overall, nf_xpatial streamlines the processing of Xenium data from platform outputs to integrated single-cell and spatial clustering datasets, providing a standardized starting point from which biologists can fine-tune parameters and proceed to hypothesis-driven spatial analyses.

18
Bacterial metagenome in plaque, saliva, and tumor samples from individuals with and without OSCC by next-generation sequencing

ERIRA, A.; ROBAYO, D. A. G.; GAMBOA, F.; CHALA, A.; MORENO, A.; ARREGUI, A. C.; MUNOZ, E.; NOGUERA, J.; TOBAR-TOSSE, F.

2026-08-29 bioinformatics 10.64898/2026.08.27.747557 medRxiv
Top 9%
0.5%
Show abstract

Background: Oral dysbiosis has been associated with oral squamous cell carcinoma (OSCC); however, most microbiome studies rely on 16S ribosomal RNA (rRNA) gene sequencing, limiting species-level taxonomic resolution. Methods: Dental plaque, saliva, and tumor tissue samples from 10 patients with OSCC and dental plaque and saliva samples from 10 healthy controls were analyzed in this exploratory cross-sectional study. DNA was extracted and subjected to shotgun metagenomic sequencing using the Illumina MiSeq platform. Sequence reads were quality filtered with fastp, taxonomically classified using Kraken2 v2.1.3, and species-level abundances were re-estimated with Bracken v2.9 following the removal of human reads and low abundance taxa. Relative abundances were compared using the Mann Whitney U test with the Benjamini Hochberg false discovery rate correction, while the Bray Curtis principal coordinate analysis was used as an exploratory approach to visualize microbial community patterns. Results: Shotgun metagenomic sequencing revealed distinct bacterial community profiles across the oral microenvironment. Dental plaque exhibited the highest taxonomic diversity and relative abundance. The control plaque was enriched in Streptococcus koreensis, Capnocytophaga sp. oral taxon 878, Treponema sp. Marseille Q4132, and Leptotrichia sp. oral taxon 498, whereas the plaque from patients with OSCC showed a higher relative abundance of Pyramidobacter piscolens, Parvimonas parva, and Gemella sanguinis. Salivary samples displayed lower diversity and a more homogeneous composition, predominantly comprising Capnocytophaga endodontalis, Prevotella jejuni, Aggregatibacter aphrophilus, and Gemella sanguinis. The tumor tissue showed relatively higher abundance of Sellimonas catena, Escherichia coli, Solobacterium moorei, and Lacrimispora sp. HJ 01. Conclusions: This exploratory study provides species-level characterization of the oral microbiome across multiple oral microenvironments in OSCC and generates hypotheses for future integrative metagenomic and functional studies investigating the potential contribution of oral bacterial communities to OSCC pathogenesis.

19
Clinical evaluation of artificial intelligence for diagnostics of antibiotic-resistant bacteria

Hessel, M.; Inda Diaz, J. S.; Sjöberg, A.; Salva-Serra, F.; Helldal, L.; Jirstrand, M.; Johnning, A.; Kristiansson, E.; Skovbjerg, S.

2026-08-31 infectious diseases 10.64898/2026.08.27.26361401 medRxiv
Top 9%
0.5%
Show abstract

Antimicrobial resistance is a public health challenge, driving the need for rapid, cost-effective diagnostic support tools. Artificial intelligence (AI) may enable prediction of susceptibility to untested antibiotics from known susceptibility results, but prospective clinical validation is required before routine use. We evaluated an AI-based decision support method, trained on invasive isolates from the European Surveillance System (TESSy), for prediction of antibiotic susceptibility in clinical Escherichia coli urine isolates. The evaluation included 99 E. coli isolates from urine samples with diversity in age, sex, and antibiotic susceptibility. Predictions were evaluated for 14 antibiotics using patient metadata and susceptibility results for 4-8 antibiotics as input. Prediction uncertainty was handled using conformal prediction, allowing abstention when confidence was insufficient. EUCAST disk diffusion test results were used as reference and genomic sequence data was used to explore mechanisms of the AI performance. Without conformal prediction, 84% of predictions were correct when susceptibility results of six antibiotics were used to predict susceptibility to eight additional antibiotics. Across all predictions generated using susceptibility results for six antibiotics as input, the major and very major error rates were 19% and 12%, respectively. Prediction errors varied between antibiotics and were associated with certain phenotypic and genotypic resistance patterns. Conformal prediction reduced errors but increased abstentions; at confidence levels of 90%, 95%, and 97.5%, the model abstained in 9.6%, 14%, and 22% of instances. The method showed promising performance, but its clinical use remains limited and may require diagnostic data beyond susceptibility test results and demographic variables.

20
Maternal cell-free RNA versus combined screening for first-trimester prediction of early-onset preeclampsia: a nested case-control study

Satorres-Perez, E.; Castillo-Marco, N.; Igual, M.; Cordero, T.; Munoz-Blat, I.; Monfort-Ortiz, R.; Marcos-Puig, B.; Simon, C.; Garrido-Gomez, T.; Perales-Marin, A.

2026-09-02 obstetrics and gynecology 10.64898/2026.08.28.26361628 medRxiv
Top 9%
0.5%
Show abstract

Background. In Europe, first-trimester combined screening with the Fetal Medicine Foundation (FMF) algorithm identifies women at increased risk of preeclampsia who may benefit from personalized aspirin prophylaxis. However, a substantial proportion of early-onset preeclampsia (EOPE) remains undetected at clinically acceptable specificity. Objective. To evaluate the first-trimester performance of MaiRa for early-onset preeclampsia (EOPE) risk stratification by benchmarking it against FMF screening in the same women, characterizing discordant patient-level classification profiles and exploring potential implementation strategies. Study Design. This secondary case-control analysis was nested within the prospective, multicentre PREMOM cohort [NCT04990141], which enrolled women with singleton pregnancies across 14 tertiary hospitals in Spain. First-trimester MaiRa and FMF risk estimates were evaluated in the same 126 pregnant women, comprising 99 uncomplicated controls and 27 EOPE cases, defined by disease onset before 34 weeks. Discrimination was compared using a stratified paired bootstrap analysis of the areas under the receiver-operating-characteristic curves. Performance was assessed at prespecified clinical thresholds, and detection rates were evaluated at fixed false-positive rates. Universal and contingent MaiRa implementation strategies were also evaluated. Results. MaiRa showed greater first-trimester discrimination for EOPE than FMF combined screening (AUC, 0.974 vs 0.900; P=.040) and consistently achieved higher detection rates across fixed false-positive rates. At false-positive rates of 5% and 10%, MaiRa detected 85.2% and 92.6% of EOPE cases, compared with 44.4% and 70.4% for FMF, respectively. Patient-level analysis demonstrated that MaiRa identified 12 of 27 EOPE cases (44.4%) classified as low risk by FMF; these pregnancies generally exhibited less abnormal conventional first-trimester profiles, including fewer maternal risk factors, lower mean arterial pressure and lower uterine artery pulsatility index, yet 8 of 12 (66.7%) subsequently developed severe EOPE. Exploratory implementation analyses showed that universal MaiRa screening achieved the highest EOPE detection, whereas a contingent strategy using FMF for triage and reflex MaiRa testing reduced molecular testing to 35.7% of pregnancies while maintaining 77.8% sensitivity and 97.0% specificity. Conclusion. MaiRa provided greater first-trimester discrimination for EOPE than conventional combined screening and detected additional pregnancies that later developed severe disease despite less abnormal conventional screening profiles. The findings suggest that maternal plasma cfRNA profiling captures biological alterations not fully reflected by combined first-trimester screening and support further prospective evaluation in an independent, unselected obstetric population. Key words: early-onset preeclampsia; first-trimester screening; cell-free RNA; liquid biopsy; Fetal Medicine Foundation algorithm; combined screening; risk stratification; aspirin prophylaxis.